NSF PAR Search | NSF Public Access Repository

HAL: Computer System for Scalable Deep Learning

https://doi.org/10.1145/3311790.3396649

Kindratenko, Volodymyr; Mu, Dawei; Zhan, Yan; Maloney, John; Hashemi, Sayed Hadi; Rabe, Benjamin; Xu, Ke; Campbell, Roy; Peng, Jian; Gropp, William (January 2020, PEARC '20: Practice and Experience in Advanced Research Computing)

We describe the design, deployment and operation of a computer system built to efficiently run deep learning frameworks. The system consists of 16 IBM POWER9 servers with 4 NVIDIA V100 GPUs each, interconnected with Mellanox EDR InfiniBand fabric, and a DDN all-flash storage array. The system is tailored towards efficient execution of the IBM Watson Machine Learning enterprise software stack that combines popular open-source deep learning frameworks. We build a custom management software stack to enable an efficient use of the system by a diverse community of users and provide guides and recipes for running deep learning workloads at scale utilizing all available GPUs. We demonstrate scaling of a PyTorch and TensorFlow based deep neural networks to produce state-of-the-art performance results.

Full Text Available

Search for: All records